Germany Black Forest Laboratory Releases Flux3 Multimodal Model: Native Audio Generation 20-Second Synchronized Video Output
Black Forest Labs unveils Flux3, a multimodal foundation model using Self-Flow architecture with dedicated encoders/decoders for image, video, audio, and motion. It unifies understanding and generation across physical and digital environments, outperforming Luma and Runway. Flux3 natively generates audio, outputs 20s sync’d video with audio, and supports text/image/video-to-video, keyframe transitions, and multilingual dialogue.....